PLOS Genetics
● Public Library of Science (PLoS)
Preprints posted in the last 7 days, ranked by how well they match PLOS Genetics's content profile, based on 862 papers previously published here. The average preprint has a 0.54% match score for this journal, so anything above that is already an above-average fit.
Mansoor, R.; Minhas, A. S.; Thomas, A.; Mansoor, A. A.; McCambridge, A. H.; Dilts, C.; Eshak, J.; Govani, D.; Nylin, B.; Trinidad, J. C.; Kanaan, A. Y.; Kara, E.; Fielder, A.; Fielder, I.; Iglendza, A.; Mukatash, Y.; Pumnea, B.; Menzel, M. M.; Shabazz-Henry, A. L.; Niepielko, M. G.; Gao, M.
Show abstract
The QxxR motif is evolutionarily conserved within DEAD-box RNA helicases, including Drosophila Me31B and human DDX6, which post-transcriptionally regulate gene expression during animal development. A pathogenic H372R substitution (QxHR to QxRR) in the QxxR motif of human DDX6 has been associated with various developmental defects, but how this motif contributes to DDX6-family protein function remains unclear. Here, we used Drosophila Me31B as an in vivo model to investigate the QxxR motifs developmental role. We generated a Drosophila strain carrying the corresponding H333R missense mutation in Me31B and characterized its effects on female fertility, embryonic viability, germline development, and Me31B-associated molecular pathways. The me31BH333R mutation reduced female fertility in a gene dose-dependent manner, with homozygous mutant females being sterile. Embryos from the mutant females also exhibited primordial germ cell defects. Despite these developmental phenotypes, the me31BH333R mutation did not significantly alter Me31B protein abundance, global ovarian transcriptome or proteome profiles, or representative germ plasm mRNA and protein localization. In contrast, bait-normalized IP-MS analysis revealed altered enrichment of selected Me31B-associated proteins, including increased association of known Me31B interactors Trailer hitch (Tral) and Ypsilon Schachtel (Yps). These findings establish Me31BH333R as an in vivo model for investigating the conserved QxxR motif and suggest that disruption of this motif compromises development not through broad changes in gene expression, but potentially through altered composition or regulation of Me31B-containing ribonucleoprotein complexes.
Pollenz, R. S.; Davenport, M.; Ruiz-Houston, K. M.
Show abstract
Phage D29 infects Mycobacterium smegmatis mc2 155 and has a non-canonical lysis cassette that encodes two endolysin proteins (Lysin A and Lysin B) and a single two transmembrane domain (TMD) protein, LysA2a similar to F1 cluster phage LysF1a. A 1TMD LysF1b homolog, LysA2b, is encoded by a gene found downstream of the tape measure. Exogenous expression of both LysA2 proteins in tandem is a cytotoxic to M. smegmatis. Deletion of lysA2a produces phages that are lysis competent with a 10-minute triggering delay and 30% plaque size reduction. Deletion of lysA2b results in severe lysis defects manifest by 70% reduced plaque size, delayed lysis timing and reduced burst size. Deletion of both lysA2 genes results in phages that are viable and show lysis phenotypes like the lysF1b deletion. Genetic complementation of lysA2b deleted phage with the lysF1b gene fully complements the lysis phenotypes but alters the triggering time to that of an F1 cluster phage. Energy poisons trigger lysis prematurely in all phages with lysA2 gene deletions. Lysis recovery mutants (LRM) isolated from phages lacking the lysA2b genes generate wild type plaque size and have point mutations that map to TMD1 or the C-terminal region of the lysA2a gene. LRMs isolated from phages lacking both lysA2 genes show premature lysis and have mutations that all map to residue C31 of a novel lipoprotein (gene 64). Deletion of gene 64 does not change wild type D29 lysis phenotypes or rescue the lysis defects of any of the lysA2 mutants. A fitness/competition assay shows that loss of the lysA2 genes imposes a substantial competitive fitness cost. These finding support a lysis regulatory network model where the 2TMD protein is maintained in an inactive state until activated by its cognate 1TMD lysis regulator and the lipoprotein has accessory function that may enhance lysis efficiency.
Kwon, H. R.; Rackley, A.; Olson, L. E.
Show abstract
Autosomal dominant gain-of-function mutations in platelet-derived growth factor receptor beta (PDGFRb) cause overgrowth of the skeleton and other connective tissue in Kosaki overgrowth syndrome. However, the target cell type and signaling pathways underlying PDGFRb-driven overgrowth are unknown. Normal postnatal growth is controlled by pituitary-secreted growth hormone (GH), which activates the STAT5 transcriptional factor to upregulate insulin-like growth factor 1 (IGF1). To investigate the role of the GH-STAT5-IGF1 pathway in PDGFRb-related overgrowth, we generated mice with a PDGFRb gain-of-function mutation in skeletal and fibroblast lineages, which resulted in STAT5 activation and gigantism. Conditional deletion of Stat5ab in connective tissue lineages rescued skeletal overgrowth and keloid-like fibrosis in the skin. Conditional deletion of GH receptor (Ghr) did not rescue overgrowth, indicating the physiological activator of STAT5 is not required for overgrowth. However, deletion of Igf1, the STAT5 target gene, and its receptor, Igf1r, in connective tissue, rescued the overgrowth phenotype. These findings demonstrate a GHR-independent STAT5-IGF1 signaling pathway in mutant connective tissue cells, which mediates PDGFRb-driven overgrowth in mice and potentially in humans with similar PDGFRB mutations.
Yarmolinsky, J.; Cavallo, F. R.; Koskeridis, F.; Yu, X.; Bouras, E.; Richenberg, G.; Costantini, I.; Ray, D.; Woolf, B.; Karhunen, V.; Ellis, L.; Haycock, P. C.; Hemani, G.; Davey Smith, G.; Tsilidis, K. K.; Zuber, V.; McKay, J. D.; Dehghan, A.; Tzoulaki, I.
Show abstract
Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4+ T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6, PDE4D, and CASP8, indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.
Saqib, M.; Chen, F.; Mistri, D. K.; Tan, L.; Wright, N.; Sarver, D. C.; Anders, R.; Aja, S.; Wong, G. W.
Show abstract
Trisomy 21 or Down syndrome (DS) affects multi-organ systems across the lifespan. The presence of an extra chromosome, along with genome dosage imbalance due to triplicated genes, contributes to the DS phenotypes. Of the DS mouse models, few are aneuploid with a freely segregating extra chromosome. We previously showed that the aneuploid Ts65Dn mice exhibit metabolic deficits consistent with the metabolic profile of DS. However, the genotype-phenotype relationships in Ts65Dn mice are complicated by the presence of triplicated genes unrelated to human chromosome 21 (Hsa21). To address this issue, we leveraged a refined model, Ts66Yah, where the extra triplicated genes in Ts65Dn have been removed. Deep phenotyping and multi-omics analyses showed that Ts66Yah mice develop pronounced and widespread metabolic disturbances. Despite sexual dimorphism in weight gain, body temperature, lipid and lipoprotein profiles, hepatic injury and adipose fibrosis, both male and female Ts66Yah mice share a common phenotype of pronounced glucose intolerance and insulin resistance, reduced mitochondrial respiratory capacity in visceral fat, altered serum inflammatory cytokine profile, and dysregulated serum and liver metabolomes. Pan-tissue transcriptomes also reveal signatures of immune activation, disrupted metabolic processes and cellular respiration, altered cytokine signaling, enhanced oxidative stress, and extracellular matrix remodeling. These combined changes across tissues disrupt metabolic homeostasis more severely in Ts66Yah than in Ts65Dn mice. Several phenotypes, including glucose intolerance, insulin resistance, tissue fibrosis, and oxidative stress were further exacerbated by an obesogenic diet. This foundational data establishes Ts66Yah as a valuable reference model for the mechanistic and comparative study of metabolic dysfunction in DS.
Efthymiou, S.; Tabata, K.; Dafsari, H. S.; Schober, E.; Latza, C.; Isaoglu, M.; Abuelrub, A.; Rad, A.; Firoozfar, Z.; Turchetti, V.; Lin, R. Q.; Maroofian, R.; Wiethoff, S.; Afzal, E.; Zafar, F.; Rana, N.; McRae, A. M.; Kaiyrzhanov, R.; Guliyeva, U.; Gulieva, S.; Melikishvili, G.; Lespinasse, J.; Vitobello, A.; Denomme-Pichon, A.-S.; Wentzensen, I. M.; Mefford, H. C.; Briere, L. C.; A Walker, M.; A High, F.; Sweetser, D. A.; Kendall, M.; Franchi, M.; Brown, M.; Latner, D.; Joset, P.; Ivanovski, I.; Alfadhel, M.; Alluhaydan, I.; Frederiksen, A. S.; Arriens, V.; Hanker, B.; Mankad, K.; Guerin, J
Show abstract
Pathogenic variants in RUBCN, encoding the Run domain Beclin-1 interacting and cysteine-rich domain-containing protein (Rubicon) have been implicated in autosomal recessive spinocerebellar ataxia 15 (SCAR15). However, the molecular mechanisms underlying disease pathogenesis remain poorly understood. Here, we report 18 individuals from 15 unrelated families harbouring biallelic RUBCN variants, who present with an aggressive neurodevelopmental disorder variably characterized by seizures, developmental delay, intellectual disability and movement abnormalities that cause regression, progressive brain atrophy and neurodegenerative features. Through functional characterization, we demonstrate that a subset of disease-associated putative truncating variants disrupt autophagy regulation. In Caenorhabditis elegans models, loss-of-function RUBCN variants result in an increased autophagic flux and impaired neuronal function, recapitulating key features in humans. Correspondingly, cellular assays reveal that nonsense and frameshift RUBCN variants lead to defective autophagy inhibition, underscoring a crucial role for RUBCN as a key negative autophagy regulator. Molecular dynamics simulations rank the eleven missense variants by structural effect, with p.Arg813Trp alone altering the target protein at both the local and the regional level and lying within the RAB7A-binding module that the truncating alleles remove altogether. Our findings establish and expand the RUBCN-related disorders as a clinically and molecularly distinct subset of autophagy-related diseases. By delineating both the genetic landscape and cellular consequences of Rubicon dysfunction, this study enhances our understanding of autophagy-related neurodevelopmental disorders and provides a foundation for future therapeutic investigations.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Trindade Pons, V.; Gillespie, N.; Smit, R. A. J.; Arias, J. D.; Yin, X.; Berndt, S. I.; Oldehinkel, A. J.; van Loo, H.
Show abstract
Obesity is a growing public health challenge, with body mass index (BMI) influenced by both genetic and environmental factors. While the role of direct genetic transmission is well established, evidence for genetic nurture effects, in which parental genotypes impact offspring through the environment, has remained mixed. This study investigates direct genetic transmission and genetic nurture effects on BMI across ages, using parent-offspring trios and pairs from the Dutch Lifelines cohort study (N = 18,897 offspring, aged 8 to 67 years). We leveraged the latest multi-ancestry BMI polygenic score (PGS) to construct transmitted (PGS-T) and non-transmitted (PGS-NT) polygenic scores, where PGS-NT consists of parental alleles not passed on to offspring and serves as a proxy for genetic nurture. Linear mixed models showed a large effect of PGS-T on offspring BMI (Beta = 0.416, p < 0.001), corresponding to a 1.85 kg/m2 increase per SD increase in PGS-T. PGS-NT had a small but significant effect (Beta = 0.026, p = 0.013), consistent with a genetic nurture effect accounting for approximately 6.6% of the effect of direct transmission. Parent-of-origin analyses showed that maternal PGS-NT effects were larger than paternal effects. PGS-T interactions with age indicated that direct transmission effects increased in childhood and stabilized in adulthood, while PGS-NT effects remained stable across age. Our findings suggest that direct genetic transmission is the dominant influence on BMI, while results are consistent with small genetic nurture effects that are driven by the maternal side.
Bresnahan, S. T.; Xiong, C.; Head, T.; Chang, Y.-H.; Bhattacharya, A.; Huang, J. Y.
Show abstract
Unmeasured confounding threatens causal inference and replicability in observational multi-omic studies across variable environments. Genetic instrumental variables (Mendelian randomization) and negative-control calibration each address complementary sources of unmeasured confounding, yet no existing framework unifies them for omics-scale mediation analysis. We introduce ICONIC, an R package that embeds genetic instruments and negative controls within a proximal causal inference framework for total-effect and mediation analysis. ICONIC implements eight estimators spanning five confounding-control strategies, supports continuous, binary, and time-to-event outcomes, and provides extensive diagnostics including sensitivity analyses that map estimator performance across plausible assumptions. Ground-truth benchmarks are calibrated to real-omics covariance structures via a hybrid generative model (GAN + feature-level Gaussian copula) rather than parametric simulation, and a companion planning tool predicts performance gains from collecting additional omic data. We demonstrate ICONIC in two case studies: identifying placental transcriptomic mediators of gestational diabetes on birth weight (n = 164), and tumor-expression mediators of smoking intensity on lung cancer survival (n = 494). Notably, ICONIC's diagnostics recommended different estimation strategies across the two scenarios, reflecting differences in the likely influence of unmeasured confounding. ICONIC is freely available at https://github.com/sbresnahan/iconic/.
Faria, S. D. S.; Bineau, J.; Moisan, R.; Legault, M.-A.; Lecluze, E.; Pincez, T.
Show abstract
The genetic risk factors of immune cytopenias are unclear. Immune cytopenias have been reported in various genetic contexts: 1) inherited error of immunity genes, mainly due to rare germline variants, 2) systemic lupus erythematosus, associated with common germline variants, 3) hematological malignancies, and 4) clonal hematopoiesis, the latter two due to somatic variants. However, the respective contribution and interaction of these variants remain to be investigated. Here, we used two large biobanks with whole genome sequencing data to systematically investigate the genetic contribution to immune cytopenia. We found that the four types of genetic variants independently contribute to immune cytopenia risk. We notably found that carriers of variants in some autosomal recessive genes of inherited error of immunity had an increased risk of immune cytopenia. Additionally, common variant-mediated risk of systemic lupus erythematosus also increased the risk of immune cytopenia. Overall, a third to a half of patients with immune cytopenia carried at least one of the four genetic risk variants investigated. Combining the four variants allowed stratifying the risk of immune cytopenia in both general and high-risk population. In general population, the 10-year incidence of immune cytopenia in the lowest and highest risk groups was 0.08% and 1.5%, respectively. In sum, this work identified that different genetic risk factors can lead to immune cytopenia. A large proportion of individuals with immune cytopenia carried an underlying genetic risk factor. Finally, combining these genetic risk factors enabled risk stratification.
Clegg, D.; Bentley-DeSousa, A.; Roczniak-Ferguson, A.; Ferguson, S. M.
Show abstract
Increased activity of leucine-rich repeat kinase 2 (LRRK2) confers Parkinson's disease risk. LRRK2 dynamically localizes to lysosomal membranes in response to various stresses, yet the mechanisms by which distinct lysosomal perturbations are communicated to LRRK2 remain unclear. Here, we show that inhibition of the lysosomal lipid kinase PIKfyve promotes LRRK2 recruitment and signaling through a pathway that requires the lysosomal chloride/proton antiporter ClC-7. ClC-7 in turn controls the accumulation of multiple Rab GTPases on lysosomes. LRRK2 signaling under these conditions requires its established Rab-binding surfaces, with Rab12 contributing significantly to this response. This pathway operates independently of CASM. In contrast, lysosomal stresses that induce CASM require both Rab-binding sites on LRRK2 and GABARAP for robust LRRK2 signaling. These findings identify ClC-7-dependent lysosomal remodeling and Rab accumulation as key features linking PIKfyve inhibition to LRRK2 signaling and reveal that distinct lysosomal stresses engage different combinations of Rab and GABARAP inputs to activate LRRK2.
Venkatesh, R.; Deo, R.; Cappola, T.; Penn Medicine BioBank, ; Ritchie, M. D.; Kim, D.
Show abstract
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and a major cause of cardioembolic stroke. Although polygenic risk scores (PRS) are well characterized to quantify inherited susceptibility for AF, they provide limited insight into the pathways and tissues underlying genetic risk, which are critical to uncover for individual risk prediction. In this study, we develop a pathway-level multi-omics representation learning framework that converts individual genetic profiles into interpretable biological features by integrating GWAS-derived pathway burden scores with tissue-specific transcriptomic pathway signals. We constructed machine learning models to assess population-level AF risk prediction performance across genomic and transcriptomic tissue contexts; the pathway-based global attention models substantially improved risk prediction performance over PRS and other baselines (AUROC improved from 0.601 to 0.738). Transformer and graph neural network frameworks then assessed individual-level pathway interpretability, revealing heterogeneous contributions from electrical signaling, cardiac development, and DNA repair pathways to AF risk. This added interpretability highlights the potential of this pathway approach to enable more mechanistically informed risk stratification than static PRS by capturing underlying heterogeneity. To independently assess whether prioritized pathways reflected cardiac regulatory biology, we compared pathway rankings with transcriptional effects predicted by the AlphaGenome foundation model. Variants in highly ranked pathways showed significantly greater predicted effects on expression in atrial and ventricular tissues (FDR = 0.032) relative to controls, providing orthogonal evidence that the model identifies biologically relevant mechanisms. Overall, this work reframes polygenic risk from a single measure of susceptibility to tissue-informed pathway mechanisms, providing a framework for interpretable genomic stratification in complex diseases.
Alquicira-Hernandez, J.; Dorans, E.; Tomofuji, Y.; Nathan, A.; Raychaudhuri, S.
Show abstract
Single-cell technologies enable linking disease-risk variants to gene regulatory effects in specific cell-state contexts. However, most so called "single-cell eQTL" studies use a "pseudobulking" strategy to identify expression Quantitative Trait Loci (eQTLs), obscuring subtle dynamic regulatory effects of disease alleles. Here, we propose Dynema (Dynamic eQTL mapping in single cells) for fast and accurate genome-wide mapping of context-dependent and independent eQTL effects at true single-cell resolution. To identify eQTLs, Dynema uses a Poisson model with cluster robust variance estimators (CRVEs) to account for correlation of single-cell profiles from the same individual. In contrast to other common methods, Dynema achieves statistical calibration and scales to genome-wide analysis in large single-cell datasets in realistic timeframes. We applied Dynema to two independent T cell datasets and identified reproducible cell-state-dependent eQTL effects. Some cell-state-dependent eQTLs are missed by pseudobulking approaches, and many others are conditionally independent from lead eQTL effects. We show that TSPAN32 and other autoimmune loci colocalize with cell-state-dependent eQTLs. Mapping context-dependent eQTLs at single-cell resolution enables the definition of the molecular effects of complex disease alleles.
Jaholkowski, P.; Parker, N.; Sveen, I. O.; Wistrom, E. D.; Fominykh, V.; Szabo, A.; Parekh, P.; Frei, O.; Smeland, O. B.; O'Connell, K. S.; Djurovic, S.; Dale, A. M.; Shadrin, A. A.; Andreassen, O. A.
Show abstract
Recent large-scale studies have enabled new knowledge about genetic underpinnings of morphological and electrophysiological alterations of the retina. Variation in retinal traits, often of neurodevelopmental origin, have been linked to major psychiatric disorders (MPDs). Here, we investigate the genetic overlap between MPDs and key retinal traits to identify underlying molecular mechanisms. We obtained genome-wide associations studies data for bipolar disorder (BD), major depression (MD), schizophrenia (SCZ), and the retinal traits retinal nerve fibre layer thickness (RNFL), ganglion cell inner plexiform layer thickness (GCIPL), and vertical cup-disc ratio (VCDR). We estimated the number of trait-influencing variants shared between traits with MiXeR and identified shared genetic loci with condFDR. Subsequently, we examined the biological pathways of the genes mapped to shared loci. This revealed that GCIPL shared the most genetic variants with MPDs (~60%), followed by RNFL (~40%), and VCDR (~20%). The genetic variants shared between retinal traits and MPDs showed disorder-specific patterns with more pronounced overlaps of SCZ and BD with RNFL, and MD negatively correlated with GCIPL. Gene-pathway analysis highlighted the importance of GABAergic neurotransmission and a two-stage neurodevelopmental process in SCZ, whereas the role of mitochondria and a weaker developmental component were observed in BD. The results also implicated synaptic functioning and gene-expression processes in MD. Furthermore, polygenic analysis suggested that the genetic architecture of retinal traits can distinguish between MPDs. Our findings indicate shared genetic underpinnings between retinal traits and SCZ, BD, and MD, implicating altered neurodevelopment and neurotransmission underlying the retinal link to major psychiatric disorders.
Kristensen, D. T.; Broendum, R. F.; Knudsen, M.; Grubach, L.; Marcher, C.; Preiss, B.; Bibi, M. L.; Hoegdall, E.; Poulsen, T.; Skov, V.; Oerskov, A. D.; Groenbaek, K.; Hansen, J. W.; Schoellkopf, C.; Cowland, J.; Andersen, M. K.; Severinsen, M. T.; Vejgaard, C.; Larsen, O. H.; Vang, S.; Boegsted, M.; Roug, A. S.
Show abstract
Large genomically annotated acute myeloid leukaemia (AML) datasets exist, but population-based contemporary cohorts remain scarce. Here we report clinicopathological, genomic, and outcome data from Danish AML patients. 2,512 AML patients were identified between 2015-2022, of whom 33.8% had available NGS data (NGS+). In patients [≤]70 years, baseline characteristics and outcomes were comparable between NGS+ and NGS- groups. In patients >70 years, more NGS+ patients received intensive treatment, but survival was similar among intensively treated patients. The distribution of mutations varied significantly by age and sex, with older age and male sex exhibiting higher frequencies of adverse-risk gene mutations. In intensively treated NGS+ patients, ELN2017 stratified 5-year OS: 58.4% (favorable), 43.4% (intermediate), and 28.2% (adverse), with hazard ratios (HRs) of 0.63 (favorable) and 1.45 (adverse) relative to intermediate. ELN2022 yielded corresponding OS rates of 56.9%, 51.8%, and 29.7%, with HRs of 0.78 and 1.86. The two models had comparable predictive performance for OS in a time-dependent model. In conclusion, outcomes of intensively treated AML patients were comparable irrespective of NGS status, underscoring the representativeness of the REFORM-AML database for the Danish AML population. Age and male sex correlated with adverse-risk mutations, and both ELN2017 and ELN2022 robustly predicted survival.
Pham, M. H.; Harvey, L. M. R.; Oliver, T. R. W.; Dunstone, E.; Lawson, A. R. J.; Nicola, P. A.; Sanghvi, R.; Hooks, Y.; Mitchell, E.; Jarman, G. L.; Wang, Y.; Abascal, F.; Jung, H.; Neville, M. D. C.; Ishida, Y.; Fowler, J. C.; Le, A. P.; Moody, S.; Marshall, H.; Brzozowska, N.; Ding, C.; Pac, C. A.; Machado, H. E.; O'Neill, L.; Latimer, C.; Humphreys, L.; Saeb-Parsy, K.; Mahbubani, K. T. A.; Baxter, J.; Rassl, D. M.; Vicario, R.; Geissmann, F.; Kabashima, K.; Bleys, R. L. A. W.; Moore, L.; Heer, R.; Coorens, T. H. H.; Behjati, S.; Hoare, M.; Campbell, P. J.; Jones, P. H.; Martincorena, I.; Ra
Show abstract
Over the course of a lifetime, somatic mutations accrue in normal human cells, causing variation in cell phenotype and engendering somatic evolution with outcomes ranging from the adaptive immune system to cancer. To inform understanding of somatic evolution in the human body we report the mutation rates and mutational signatures of 53 normal cell types. Most show evidence of linear mutation accumulation over time with single base substitution mutation rates ranging from ~3.5/year/diploid genome in spermatogonia and sperm, to ~20/year in postmitotic neurons, ~50/year in mitotically active colorectal epithelial cells, ~60/year in kidney proximal tubule cells and hepatocytes, 100s/year in sun-exposed skin epidermal cells and 10-50/year in the remainder. Certain cell types, including skin epidermis, cardiac myocytes, bladder urothelium, kidney proximal tubule cells, and hepatocytes, show substantial variability in mutation burdens around the linear age trend, indicating the influence of additional factors which differ between individuals and modulate mutation accumulation, including exogenous mutagen exposures. At least 18 single-base substitution and nine small insertion and deletion mutational signatures are present, some in all cell types, some in a subset and others in a single cell type. Known exogenous mutagen exposures and endogenous mutational processes account for some mutational signatures, but the origins and mechanisms underlying many are uncertain. This comprehensive survey of mutagenesis provides a foundation for understanding somatic evolution of human cell populations in health and disease.
Bourne, N. G.; Payne, L.; Manzi, S.; Besnard, G.; Vorontsova, M. S.; Jobson, R. W.; Chomicki, G. S.; Dunning, L. T.
Show abstract
Determining the correct donor species/lineages of grass-to-grass lateral gene transfer (LGT) is vital for deducing specific donor features that could help inform the mechanism of transfer. This requires a dataset spanning a broad range of species to achieve the phylogenetic resolution necessary for precise donor inference. As grass-to-grass LGT often involves the transfer of multi-gene DNA fragments, they can contain additional sequences that allow for accurate orthologous comparisons, such as nuclear DNA of plastid origin (NUPTs). Here we systematically scan for NUPTs in the genomes of four Alloteropsis semialata accessions, whose LGTs have previously been characterised. Using the abundant Panicoideae chloroplast sequences, we reconstruct NUPT phylogenies and infer two lateral acquisitions: one from Paniceae/Digitaria and another from Andropogoneae/Eremochloa adjacent to a previously identified LGT. We then assembled and included an additional 12 Eremochloa chloroplast genomes in the analysis and showed the likely donor was Eremochloa attenuata. Subsequent short-read mapping from E. attenuata to the nuclear region flanking this NUPT showed consistent coverage across the region, including the previously identified LGT, supporting co-transfer. Overall this study highlights the potential for NUPTs to better identify the donors of grass-to-grass LGT.
Dos Santos, M.; Ohtsuki, H.; Mullon, C.
Show abstract
Reputation plays a major role in supporting cooperation among unrelated individuals through indirect reciprocity. By helping others, individuals build a good personal reputation and receive greater benefits from future partners. Most models of indirect reciprocity assume that a person's reputation reflects only their own behaviour. Yet in many societies, people are also judged by their family's reputation. How family reputation affects the evolution of cooperation, and whether reliance on it can itself evolve, remain unclear. Here we show that reputation inheritance expands the conditions under which indirect reciprocity favours cooperation, increasing helping and favouring greater reciprocity. Greater reciprocity in turn favours stronger reliance on inherited reputation, creating a positive feedback that stabilises cooperation, especially when interactions are infrequent or personal behaviour is difficult to observe. This feedback arises because cooperation generates future benefits both for the individual, through their personal reputation, and for their descendants, through inherited reputation. Reputation inheritance thereby provides a route via which kin selection and reciprocity, often treated as alternative explanations for cooperation, can reinforce one another. Our model helps explain why family-based reputation occurs across diverse human societies and provides an evolutionary framework for studying phenomena organised around family standing, including kin-based institutions, feuds between families and honour-based violence within them.
Pham, K.; Nicastro, G. G.; Long, A. R.; Aravind, L.; Wilke, C. O.; de Souza, R. F.; Bayer-Santos, E.
Show abstract
Microorganisms across all domains of life engage in molecular conflict, deploying toxins to inhibit competitors or respond to biological threats. Among these, ribonuclease toxins are particularly widespread and diverse. A substantial fraction is associated with the BECR fold, a compact /{beta} architecture that supports RNase activity despite extensive divergence. Although several canonical members are well characterized, many BECR-fold proteins remain difficult to identify because of low sequence similarity, variation in catalytic residues, and structural elaborations that obscure evolutionary relationships. The growing availability of high-confidence protein structure predictions provides an opportunity to reassess this deeply divergent protein landscape. Here, we integrate iterative profile-HMM searches, profile-similarity networks, structural analyses, active-site mapping, and genomic context to examine BECR proteins across the tree of life. Our analysis resolves an expanded BECR-fold landscape comprising canonical BECR and BECR-like superfamilies, refines the organization of canonical BECR proteins and identifies previously unrecognized families. We further validate BECR-Tox2 as a toxin neutralized by a cognate immunity protein and show that its homologs occur in both Menshen-like anti-phage systems and polymorphic toxin loci. Together, these findings expand and clarify the BECR-fold landscape and provide a framework for identifying and interpreting highly divergent proteins of this fold.
Bowness, J. S.; Bernal Martinez, A.; Barinka, J.; Schulte-Schrepping, J.; Renders, S.; Waclawiczek, A.; Leppa, A.-M.; Trumpp, A.; Raffel, S.; Haas, S.; Velten, L.
Show abstract
To sustain blood formation, hematopoietic stem and progenitor cells (HSPCs) coordinate a multitude of cell biological processes, from cell cycle control and stress responses to lineage priming. While many genetic regulators of high-level HSPC function have been identified, how HSPCs coordinate more basal cell biological programs, and how such programs relate to stem cell function, remains incompletely understood. Here we use Perturb-seq to profile the transcriptional consequences of targeting 520 genes by CRISPRi in primary mouse HSPC cultures. We developed an analytical strategy to separate perturbation-induced changes in cell-state abundance and clonal heterogeneity from cell-state-local transcriptional effects. From these local perturbation signatures, we identified 19 gene regulatory programs (GRPs) that are defined by co-regulation in response to genetic perturbation, in contrast to co-expression or human curation, and align well with cell biological processes. By decomposing gene expression data from functional and clinical studies into program activity, we show that GRP activities associate with, and predict, phenotypes such as clonal output after transplantation, as well as survival and drug response in retrospective acute myeloid leukemia (AML) cohorts. Together, our study establishes perturbation-derived co-regulation programs as an interpretable framework for linking genetic regulators, cell-biological processes and stem-cell-associated phenotypes.